Nvidia2026-07-06 03:29:56Nvidia GB300 NVL72 Hits 61,400 Concurrent AI Agents per Megawatt in AA-AgentPerf BenchmarkNvidia said its GB300 NVL72 system delivered 61,400 concurrent AI agents per megawatt in the new AA-AgentPerf benchmark from independent evaluator Artificial Analysis, compared with roughly 2,600 for the previous-generation H200. Under the benchmark’s 20-token-per-second and 60-token-per-second service tiers, Nvidia said the GB300 maintained about a 20x advantage on the key “agents per megawatt” metric, while per-GPU density reached 57.5 agents versus 1.4 on H200, implying about a 40x gap on that measure. The significance of the result is not only the hardware gain, but also the benchmark’s design. AA-AgentPerf is positioned as a benchmark built specifically for AI agents rather than single-shot inference. It replays real coding-agent trajectories, includes long context windows, tool interactions, and chained model calls, and evaluates systems under service-level objectives such as output speed and time-to-first-token. The shift suggests AI inference evaluation is moving away from token throughput alone toward system-level efficiency, service density, and the number of agents a power budget can sustain.510